文章背景与核心概要
在具身智能和机器人学领域,构建能够进行有效前瞻规划的世界模型一直是核心挑战之一。传统的生成式世界模型往往依赖于计算成本高昂且容易产生累积误差的像素重建。近年来,联合嵌入预测架构(JEPA)虽然允许模型在不进行像素重建的情况下直接针对视觉指定的目标进行规划,但单纯的潜在空间预测常面临表征崩溃或丢失控制相关关键信息的风险。
为了克服这些局限性,来自北京通用灵动人工智能(GENISOM AI)的研究团队提出了一种端到端的 JEPA 世界模型,专为目标条件化机器人规划设计。该方法在潜在空间预测的基础上创新性地引入了两个关键组件:一是逆动力学模型(IDM),用于防止潜在空间坍缩,并确保潜在状态转移能够真实反映生成这些转移的动作;二是状态Alignment(SA),用于将连续的表征与其物理构型和运动进行接地绑定。
实验结果表明,该模型在四个基准测试任务中表现卓越,在 TwoRoom(100%)、PushT(98%)和 OGBench-Cube(87%)上均取得了顶尖的成功率,在 Reacher 任务上的表现也与基线模型不相上下。消融实验进一步证实,结合状态对齐能够持续提升规划性能,超越仅使用逆动力学模型的效果,为具身智能的可扩展物理世界模型开发提供了重要参考。
Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning
Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning
Authors: Muyuan Liu (1), Yue Huang (1), Zheng Liang (1), Xiang Gao (1)
(1) GENISOM AI, Beijing, China
Authors: Muyuan Liu (1), Yue Huang (1), Zheng Liang (1), Xiang Gao (1)
(1) GENISOM AI, Beijing, China
📌 Summary
📌 Summary
This paper introduces an end-to-end Joint-Embedding Predictive Architecture (JEPA) world model designed for goal-conditioned robotic planning. While action-conditioned JEPA models allow planning toward visually specified goals without pixel reconstruction, latent prediction alone often fails to ensure that the learned representations retain control-relevant information.
This paper introduces an end-to-end Joint-Embedding Predictive Architecture (JEPA) world model designed for goal-conditioned robotic planning. While action-conditioned JEPA models allow planning toward visually specified goals without pixel reconstruction, latent prediction alone often fails to ensure that the learned representations retain control-relevant information.
To overcome this limitation, the authors augment latent prediction with two key components: 1. Inverse Dynamics (IDM): Prevents latent collapse and makes latent transitions informative of the actions that generated them. 2. State Alignment (SA): Grounds consecutive representations in their physical configurations and motions.
To overcome this limitation, the authors augment latent prediction with two key components: 1. Inverse Dynamics (IDM): Prevents latent collapse and makes latent transitions informative of the actions that generated them. 2. State Alignment (SA): Grounds consecutive representations in their physical configurations and motions.
Across four benchmark tasks, the proposed model excels—achieving top-tier success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to baseline models on Reacher. Ablation studies further demonstrate that integrating state alignment consistently enhances planning performance beyond using inverse dynamics alone.
Across four benchmark tasks, the proposed model excels—achieving top-tier success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to baseline models on Reacher. Ablation studies further demonstrate that integrating state alignment consistently enhances planning performance beyond using inverse dynamics alone.
📋 Document Details
📋 Document Details
- arXiv ID: arXiv:2609.03565 [cs.RO]
- Primary Subject: Robotics (
cs.RO) - Secondary Subjects: Artificial Intelligence (
cs.AI), Machine Learning (cs.LG) - Submission Date: September 3, 2026
- Conference Acceptance: Accepted to the IROS 2026 Workshop on Physical World Models for Scaling Embodied AI (PWMS 2026)
- Paper Metrics: 5 pages, 4 figures, 2 tables
- arXiv ID: arXiv:2609.03565 [cs.RO]
- Primary Subject: Robotics (
cs.RO)- Secondary Subjects: Artificial Intelligence (
cs.AI), Machine Learning (cs.LG)- Submission Date: September 3, 2026
- Conference Acceptance: Accepted to the IROS 2026 Workshop on Physical World Models for Scaling Embodied AI (PWMS 2026)
- Paper Metrics: 5 pages, 4 figures, 2 tables
🔗 Full-Text & Resources
🔗 Full-Text & Resources